Papers with LLM-based approach
A Closer Look at Claim Decomposition (2024.starsem-1)
Copied to clipboard
| Challenge: | Recent work uses claim decomposition to determine how well supported a claim is for applications in factual precision of generated text, entailment of human generated text and claim verification. |
| Approach: | They propose an LLM-based approach to generating decompositions inspired by Bertrand Russell’s theory of logical atomism and neo-Davidsonian semantics and demonstrate its improved decomposing quality over previous methods. |
| Outcome: | The proposed method improves on the FActScore and a Bertrand Russell-inspired approach to generating decompositions inspired by neo-Davidsonian semantics and improves decomposability quality. |
“Stupid robot, I want to speak to a human!” User Frustration Detection in Task-Oriented Dialog Systems (2025.coling-industry)
Copied to clipboard
Mireia Hernandez Caralt, Ivan Sekulic, Filip Carevic, Nghia Khau, Diana Nicoleta Popa, Bruna Guedes, Victor Guimaraes, Zeyu Yang, Andre Manso, Meghana Reddy, Paolo Rosso, Roland Mathis
| Challenge: | Detecting user frustration in task-oriented dialog systems is imperative for maintaining overall user satisfaction, engagement and retention. |
| Approach: | They compare out-of-the-box methods for user frustration detection with open-source methods . they find an LLM-based approach is promising, as it captures both emotion and dialog breakdowns . |
| Outcome: | The proposed method outperforms open-source methods in detecting user frustration in a TOD system. |
EduPulse: A Practical LLM-Enhanced Opinion Mining System for Vietnamese Student Feedback in Educational Platforms (2026.eacl-industry)
Copied to clipboard
| Challenge: | EduPulse is a system designed specifically to analyze student feedback in Vietnamese. |
| Approach: | They propose a system that analyzes student feedback in Vietnamese to improve opinion mining. |
| Outcome: | The proposed system performs four opinion analysis tasks in Vietnamese . it is scalable and maintainable, and it is cost-effective, the authors show . |
Efficient Out-of-Scope Detection in Dialogue Systems via Uncertainty-Driven LLM Routing (2025.acl-industry)
Copied to clipboard
| Challenge: | Out-of-scope (OOS) intent detection is critical in task-oriented dialogue systems . without effective OOS detection, such inputs could lead to incorrect responses, reduced user trust, and eventual system failures. |
| Approach: | They propose a modular framework that combines uncertainty modeling with fine-tuned large language models (LLMs) their method yields state-of-the-art results on key OOS detection benchmarks . |
| Outcome: | The proposed framework yields state-of-the-art results on key OOS detection benchmarks including real-world OOS data. |
Coding Open-Ended Responses using Pseudo Response Generation by Large Language Models (2024.naacl-srw)
Copied to clipboard
| Challenge: | Existing pipelines for survey research using open-ended responses require time and cost-consuming manual tasks. |
| Approach: | They propose an LLM-based method to automate parts of the grounded theory approach . they generate and annotate pseudo open-ended responses and use them as training data . |
| Outcome: | The proposed method is highly efficient andcost-saving compared to human-based methods. |
Transforming Podcast Preview Generation: From Expert Models to LLM-Based Systems (2025.acl-industry)
Copied to clipboard
| Challenge: | Podcasts, videos, and other long-form talk content requires significant time investment to assess their relevance. |
| Approach: | They propose an LLM-based approach for generating podcast episode previews and deploy it at scale, serving hundreds of thousands of podcast previews in a real-world application. |
| Outcome: | The proposed approach outperforms a baseline built on top of various ML expert models and offers a 4.6% increase in user engagement with preview content and a 5x boost in processing efficiency. |
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms (2025.acl-long)
Copied to clipboard
| Challenge: | Social media platforms use machine learning and artificial intelligence to maximize user engagement, but can indirectly cause exposure to harmful content. |
| Approach: | They propose a re-ranking approach using Large Language Models to assess and rerank content sequences using large annotated data sets. |
| Outcome: | The proposed method significantly outperforms existing proprietary moderation methods on three datasets, three models and across three configurations. |
Harnessing LLMs for Temporal Data - A Study on Explainable Financial Time Series Forecasting (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Recent advances in machine learning and artificial intelligence have opened up numerous opportunities and challenges in financial time series forecasting. |
| Approach: | They propose to use Large Language Models for explainable financial time series forecasting to leverage cross-sequence information and extract insights from text and price time series. |
| Outcome: | The proposed model outperforms ARMA-GARCH and gradient-boosting tree models while underperforming on other models. |
Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | In traditional face-to-face therapy, the assessment of therapeutic alliance is not directly translated to text-based settings. |
| Approach: | They propose an automatic approach to understand the development of therapeutic alliance in text-based counseling by using large language models. |
| Outcome: | The proposed approach demonstrates that the framework is effective in identifying the therapeutic alliance in text-based counseling. |
How to Contextualize Empirical Data for Risk Analysis with LLMs: A Case Study of Power Outages (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly being considered for high-stakes decision-making, yet their application in statistical risk analysis remains largely underexplored. |
| Approach: | They propose a method for extracting key information from raw data and translating it into structured contextual input within the LLM prompt. |
| Outcome: | The proposed approach significantly improves the LLM’s performance in risk assessment tasks. |
Dialogue Summarization with Mixture of Experts based on Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies for dialogue summarization use one model at a time or treat it as a black box. |
| Approach: | They propose an LLM-based approach with role-oriented routing and fusion generation to utilize mixture of experts for dialogue summarization. |
| Outcome: | The proposed approach produces informative and accurate dialogue summarization on widely used datasets. |
Learning Multimodal Contrast with Cross-modal Memory and Reinforced Contrast Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Using a memory module, we learn multimodal contrast using encoding-decoding paradigm . multimodal information are used in many applications, including news feeding, social media, etc. |
| Approach: | They propose an LLM-based approach for learning multimodal contrast following the encoding-decoding paradigm . they use a memory module with reinforced contrast recognition to enhance learning . |
| Outcome: | The proposed approach outperforms baseline and state-of-the-art studies on four English and Chinese benchmark datasets. |
Eliciting Motivational Interviewing Skill Codes in Psychotherapy with LLMs: A Bilingual Dataset and Analytical Study (2024.lrec-main)
Copied to clipboard
Xin Sun, Jiahuan Pei, Jan de Wit, Mohammad Aliannejadi, Emiel Krahmer, Jos T.P. Dobber, Jos A. Bosch
| Challenge: | Motivational interviewing (MI) is an essential, directive, client-centered counseling technique. |
| Approach: | They propose a bilingual dataset of MI conversations in English and Dutch . they propose an approach to elicit MISC expertise from Large language models . |
| Outcome: | The proposed approach yields results aligned with expert annotations and maintains consistent performance across languages. |
Large Language Model-Based Event Relation Extraction with Rationales (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for ERE rely on large language models, but they face limitations. |
| Approach: | They propose an LLM-based approach with rationales for the ERE task . LLMERE transforms ERE into a question-and-answer task that may have multiple answers . |
| Outcome: | Experimental results show that LLMERE improves over existing methods. |
TempCompass: Do Video LLMs Really Understand Videos? (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks on video large language models lack a comprehensive feedback on temporal perception ability . current models cannot distinguish between different temporal aspects and are limited in task formats . |
| Approach: | They propose a benchmark to evaluate temporal perception ability of video large language models . they construct conflicting videos that share the same static content but differ in a specific temporal aspect . |
| Outcome: | The proposed benchmarks show that video large language models exhibit poor temporal perception ability. |
Measuring Contextual Informativeness in Child-Directed Text (2025.coling-main)
Copied to clipboard
Maria R. Valentini, Téa Y. Wright, Ali Marashian, Jennifer M. Ellis, Eliana Colunga, Katharina von der Wense
| Challenge: | Recent advances in natural language processing (NLP) have made it possible to generate children's stories with a single word. |
| Approach: | They propose a task of measuring contextual informativeness in children's stories and a large language model to automate the task. |
| Outcome: | The proposed method outperforms baselines and can generalize to measuring contextual informativeness in adult-directed text. |
Large Language Models Are Natural Video Popularity Predictors (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can better capture cultural and social factors such as viewing intensity and geographic spread of video content. |
| Approach: | They propose to use Large Language Models to capture cultural and social factors that influence video popularity and generate interpretable, attribute-based explanations. |
| Outcome: | The proposed model captures both engagement intensity and geographic spread on 13,639 popular videos, while the neural network's predictions reach 82% without fine-tuning. |
It is not a piece of cake for GPT: Explaining Textual Entailment Recognition in the presence of Figurative Language (2025.coling-main)
Copied to clipboard
| Challenge: | Figure-based language is used to convey opinions, ideas, or emotions in texts and dialogues. |
| Approach: | They evaluate the capabilities of Large Language Models to address TER and generate textual explanations of TER predictions. |
| Outcome: | The proposed model outperforms the open-source models in Zero- and Few-Shot Learning settings and shows significant performance improvements. |
Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments (2025.acl-long)
Copied to clipboard
| Challenge: | Specialized non-LLM NLP-based solutions lack reasoning and are not able to infer values not explicitly present in documents. |
| Approach: | They propose a novel LLM-based approach that organizes VRDs into localized semantic textual segments called semantic blocks. |
| Outcome: | The proposed approach outperforms the state-of-the-art on public VRD benchmarks by 1-3% in F1 scores and is resilient to document formats previously not encountered. |
Computational Analysis of Conversation Dynamics through Participant Responsivity (2025.emnlp-main)
Copied to clipboard
| Challenge: | Growing literature explores toxicity and polarization in discourse, with comparatively little work on characterizing what makes dialogue prosocial and constructive. |
| Approach: | They develop and evaluate methods for quantifying responsivity through semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
| Outcome: | The proposed method is based on semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |